In the previous articles on understanding protein structure, we saw that the region where two amino acids form a peptide bond creates a rigid, flat plane with very little rotation because of resonance, while the single bonds on either side that connect this plane to the central alpha carbon (Cα) of an amino acid can rotate. These two rotatable angles, called the φ (phi) and ψ (psi) dihedral angles, can be thought of simply as the basic rotational angles that determine how the protein backbone can bend and twist. As different combinations of these two dihedral angles are formed, the protein backbone can therefore adopt a wide variety of three-dimensional structures. In theory, an enormous number of φ and ψ combinations seem possible. In an actual protein, however, the allowed combinations are considerably restricted because steric clashes occur when atoms come too close to one another. Ramachandran calculated these restrictions one by one and represented them as a map known as the Ramachandran plot. Anfinsen then showed experimentally that even when a protein was completely unfolded and then left to refold, it returned to the same structure, and argued that the information determining how a protein folds is already contained in its amino acid sequence. Levinthal likewise showed through calculation that if a protein had to explore every theoretically possible structure one by one before finding its native structure, it would be impossible to explain how real proteins fold within such a short time.
These historical questions surrounding protein structure later developed into a new perspective known as energy landscape theory, influenced by thermodynamics and statistical physics. This theory explains, in terms of free energy and probability, how a protein folds towards a particular native structure among the enormous number of structures it could potentially adopt. In particular, the folding funnel, which illustrates how a protein moves among many possible structures while being biased overall towards lower free energy and folding rapidly, has become one of the most representative ways of depicting protein folding. But is this funnel merely an intuitive picture designed to make the idea easier to understand? Or is it a theoretical framework that can physically explain the actual movements of proteins, the probabilities of different states, and the transitions between them? In this article, we will look more closely at the ideas from which energy landscape theory developed, and then follow how this theory connects to actual calculations through Folding@home and molecular dynamics simulations.
Background to energy landscape theory
First, let us look at the background of how energy landscape theory developed. From the late 1980s, studies began to appear that viewed protein folding not as a process following one predetermined pathway, but as a process occurring probabilistically on a free-energy landscape. Concepts such as the rugged landscape, minimal frustration, and the funnel-like landscape developed through a number of studies during this period, with Peter Wolynes, José Onuchic, and others playing central roles in developing these ideas into a unified theory. Then, in 1997, Onuchic, Luthey-Schulten, Wolynes, and colleagues published a review bringing together the research findings and concepts that had accumulated up to that point. This was “Theory of Protein Folding: The Energy Landscape Perspective”, published in the Annual Review of Physical Chemistry. Regarded as one of the representative papers that most systematically organised energy landscape theory, it is often considered a standard account of modern protein folding theory and provided a framework for redefining protein folding as a problem of stochastic dynamics. In particular, the funnel model introduced in this paper has become one of the most representative conceptual diagrams used today to explain protein folding and is widely used in fields such as biochemistry and structural biology.
Put simply, the central message of this paper is that a protein does not fold along one predetermined pathway. Instead, it explores many possible pathways, undergoing biased diffusion towards lower free energy until it eventually reaches the stable native region, the native basin. As the abstract of the paper states, “The energy landscape theory of protein folding is a statistical description of a protein's potential surface.” In other words, energy landscape theory does not view one particular protein structure from a deterministic perspective, but describes all possible structures probabilistically through statistical mechanics. The important point here is that protein folding is not thought of as a process of moving straight down towards one structure. Even after folding, a protein continues to fluctuate and move under the influence of the thermal motion of the surrounding water molecules. For this reason, the native structure that we commonly think of as the completed folded state can be understood not as one completely fixed structure, but as a native ensemble consisting of many very similar structures.
This view of proteins as existing as a probability distribution of multiple structures is consistent with the modern structural biology view that proteins remain dynamic entities, continually moving across the energy landscape even after they have folded. It also makes it possible to describe, within the statistical-mechanical framework of energy landscape theory, phenomena such as intrinsically disordered regions (IDRs), which have recently attracted enormous interest in biology, and liquid-liquid phase separation (LLPS) driven by these IDRs.
Thermodynamics and Kinetics
According to this paper, a protein must satisfy two crucial conditions as it folds: it must be thermodynamically foldable and it must be able to fold very quickly, or be kinetically foldable. Being thermodynamically foldable means that the protein is sufficiently stable in its folded state and that the native ensemble forms the most stable native basin on the free-energy landscape. Being kinetically foldable, on the other hand, means that on the way to reaching the native state, the protein does not remain trapped for too long in kinetic traps or behind high energy barriers, but can complete folding within a biologically meaningful short period of time. For a protein that needs to fold into a stable structure and function, both conditions are important. Together, these two concepts explain within a single picture both Anfinsen’s question of where a protein folds towards and Levinthal’s question of how, given so many possible structures, it can reach that structure so quickly. As mentioned earlier, most of the leading figures behind this theory came not from biology but from physics, particularly statistical physics. Energy landscape theory is therefore not a theory that explains the final structure of a protein, but a physical framework for explaining how a protein explores the space of possible structures.
Key concepts of energy landscape theory
To understand in more detail the representative funnel model that best captures the central idea of this paper, let us look one by one at the concepts and terms that frequently appear in this paper and in energy landscape theory more generally.
Native state
The native state does not simply mean one completely fixed structure, like a single photograph of a protein. It is more accurate to understand it as a collection of very similar structures capable of carrying out their function under physiological conditions—in other words, a native ensemble. On the energy landscape, the region in which these stable ensemble structures are gathered is called the native basin, and in the funnel model it is represented as the deepest valley.
Native contact
A native contact is a non-covalent contact formed between particular amino acid residues that lie close to one another in space in the native structure. These contacts can include hydrogen bonds, hydrophobic packing, salt bridges, and van der Waals interactions. As these native contacts gradually form during folding, the protein moves towards the stable native basin.
Molten globule state
This is a representative intermediate state that can appear during protein folding. Its overall size and shape are similar to those of the native state, and much of the secondary structure and hydrophobic core has already formed, so from the outside it can look as though the protein is already quite well folded. Inside, however, the organisation is not yet complete. Detailed packing, particularly the arrangement of side chains and the precise organisation of the tertiary structure, has not yet fully formed. For reference, the term molten globule literally conveys the image of a rounded, compact shape like a small droplet or ball—a globule—whose interior is not rigidly fixed but remains loose and fluid-like, or molten.
Discrete traps and glass transition
An energy landscape is not always smooth. As a protein moves towards the native basin, there may be small valleys with lower free energy than their surroundings. Because these valleys are relatively stable, the protein can easily remain in them for a while. These small valleys or depressions are called traps. The lowest point within such a depression is called a local minimum, while the entire surrounding region is called a basin. In other words, it is better to think of a trap not simply as a single point, but as a local basin in which a protein may move back and forth among several similar structures yet remain for some time without easily escaping. If thermal fluctuations are sufficient, the protein can cross an energy barrier and escape from this kinetic trap. However, if there are many such traps and they are deep enough to hold the protein for long periods, structural exploration slows down and the overall folding process becomes much more complicated. If this slowing becomes very severe, structural exploration may become so slow that it appears frozen or almost stopped, a situation sometimes described using physical concepts such as a glassy state or glass transition. In other words, the more rugged the landscape is, the more likely the protein is to become trapped in these depressions, and the slower folding becomes.
Stability gap
The stability gap, or energy gap, can be understood as the difference in free energy between the native basin and nearby misfolded states or molten globule states. If the free energy of the native basin is considerably lower than that of the surrounding states, the protein will show a stronger preference for the native state. Therefore, the larger the energy gap, the greater the thermodynamic stability separating the native state from other states, and the more strongly the protein is drawn towards the energetically more favourable native state.
Ensemble
In this paper, the term ensemble is connected to its original meaning in statistical mechanics. An ensemble is a collection of the many microstates that a protein can adopt under the same conditions, for example at a constant temperature and pressure. Therefore, an unfolded protein is not one single structure. There are many different possible structures, and the protein continually moves among them. The same is true of the folded state. Rather than being one completely fixed structure, the native state can also be understood as a native ensemble made up of many very similar structures. Thus, on an energy landscape, the low-free-energy region around the native state is better thought of not as one exact point, but as a basin in which similar structures are gathered and the protein continually moves among smaller valleys within this stable region. Ultimately, energy landscape theory explains the movements and probabilities among structures across this entire collection of conformations.
One point I would particularly like to add here regarding protein ensembles is that, even at the time this paper was published, the concept of the ensemble already formed an important part of the background to energy landscape theory, but it was not yet emphasised as strongly as it is today in understanding protein structure and function. As research into protein dynamics and structural diversity developed, the view that the native state should be understood not as one fixed structure, but as an ensemble in which many very similar structures continually interconvert, became much more important, in contrast to the common assumption that once a protein folds, it becomes fixed in a single structure. This collection of multiple structures is called the native ensemble.
If we consider that proteins inside a cell at 37°C are continually colliding with water molecules and vibrating, this is actually quite easy to understand. In the midst of constant vibration and thermal motion, atoms are not completely fixed in a single structure, but continually undergo small changes in position and structure. Enzymes, too, repeatedly open and close and can move among various intermediate states. The native state, therefore, can be thought of not as one fixed point with the lowest free energy, but as a region in which the protein continually moves among several similar structures even within one stable well. Inside the cell, the environment can continually change in response to factors such as pH, temperature, and phosphorylation, and the energy landscape can change accordingly. As a result, under some conditions, a different well may become the more stable state. I would like to return to this in more detail in a future article on phase transitions, a subject in which I am personally very interested.
Non-native interactions and nonspecific hydrogen bonding
The more interactions there are that help a protein form its native structure, the more the overall energy landscape takes on a funnel shape leading towards the native basin. Conversely, problems arise if there are too many interactions that are unrelated to the native structure. For example, a nonspecific hydrogen bond that happens to form during folding may temporarily stabilise a particular structure. This can cause the protein to deviate from the path it had been following and fall into a local trap, and the more numerous these nonspecific interactions become, the more rugged the energy landscape becomes. A good energy landscape can therefore be described as having evolved towards a minimally frustrated landscape, in which energetic frustration is minimised so that the native structure is stabilised while incorrect structures do not become excessively stable.
A representative example of such interactions is nonspecific hydrogen bonding. This refers to hydrogen bonds that happen to form during folding apart from the specific hydrogen bonds required to form the native structure, such as the \(i \rightarrow i+4\) hydrogen bonds in an α-helix. Interestingly, these nonspecific interactions also play important roles in phenomena such as intrinsically disordered regions (IDRs) and liquid-liquid phase separation. I would like to return to this in future related articles.
Thermal fluctuations inside the cell
Proteins fold in the watery environment inside living cells. In reality, however, the inside of a cell is a complex environment containing not only water but also very high concentrations of proteins, nucleic acids, ions, metabolites, and many other molecules. All of these are in constant thermal motion, so the atoms within a protein are also continually colliding with the surrounding environment, including water, exchanging energy and fluctuating as they do so. These continuous, tiny movements are called thermal fluctuations. As folding proceeds, amino acid residues that were originally far apart in three-dimensional space may come closer together, creating new contacts—in other words, new interactions. These interactions are not fixed. Because of thermal fluctuations, a contact may form briefly and then break apart again, while residues that were previously separated may approach each other and form new contacts. Such small changes can affect the movements and interactions of other residues and can lead to a chain of structural changes. Protein folding is therefore not a smooth process that moves in one direction from beginning to end. It proceeds through countless small structural changes and transitions, while gradually becoming biased overall towards lower free energy. In this sense, it can be described as a stochastic process. This is why, on an energy landscape, a protein is compared to a ball rolling down a mountain. The ball does not roll in a straight line from the top of the mountain directly to the bottom of the valley. On the way down, it may temporarily fall into a small depression, climb back out, and move into another valley. Yet the overall movement still has a direction. Likewise, even though a protein continues to undergo local fluctuations because of thermal motion, overall it eventually moves towards lower free energy and closer to the native state.
Now, based on these concepts, let us look at protein folding as represented by the funnel shape. The image below has been reconstructed to closely resemble the image presented in the original paper.
The wide opening at the top of the funnel, the unfolded ensemble, corresponds to the completely unfolded state of the protein. Here, the protein does not exist as a single completely unfolded structure. Instead, there are many different possible structures. This region therefore contains a great variety of structural states and has a high degree of freedom to move in many different directions. Here, an ensemble means the countless random microstructures that can exist under the same environmental conditions, and this also means that entropy—which reflects the number of microstates capable of making up a particular macrostate—is very high.
As the protein begins to fold, several structural changes occur at the same time and the protein gradually becomes more compact. The range of possible structures gradually narrows. During folding, substantial secondary structure can form and a hydrophobic core can develop, giving rise to intermediate states such as a molten globule. Even during this process, however, the protein continues to move constantly because of thermal fluctuations, and its structure keeps changing. Some conformations move downward as native contacts form, while others may remain temporarily in local traps because of incorrect contacts. As folding progresses further, the contacts and overall structural arrangement that make up the native structure gradually become established, and the direction towards the native state becomes increasingly clear.
The walls of the funnel are not smooth but rugged. This is because interactions unrelated to the native structure can form during folding, creating small depressions. For example, if a nonspecific hydrogen bond happens to form between an inappropriate N–H and O=C pair and temporarily stabilises a particular structure, the protein may fall into that structure and remain there for a while. However, the protein can escape from such kinetic traps, which affect the rate of folding, through thermal fluctuations and go on to explore other structures. The final destination of folding is the native state, the global minimum located at the very bottom of the funnel. This region, often compared to the deepest well on the energy landscape, is where the protein can exist most stably under physiological conditions. The larger the energy gap between the native state and other structures, the more stable the native state becomes relative to them, and the more strongly the protein will favour it.
From theory to reality: the Folding@home project
After following the idea of the energy landscape this far, a natural question arises. Do all these things described by the energy landscape actually happen in real proteins? Does a protein really move among countless microstates while tending towards lower free energy? Is the funnel model we have been looking at the result of observing the actual movements of proteins, or is it simply a beautiful theoretical picture based on thermodynamics and statistics?
To answer this question, we need to return to the starting point of statistical mechanics. As we saw in the articles on thermodynamics, the macroscopic phenomena we observe in nature appear as the statistical results of countless microscopic events that we cannot see. Thermodynamic quantities such as temperature, pressure, and entropy are likewise the result of averaging and statistically describing the motions of vast numbers of atoms and molecules. Proteins are no exception. Even while a single protein folds, countless atoms are continually moving, forming interactions, separating, and coming together again. On top of this, the surrounding water molecules and ions jostle with the protein and exchange energy with it. Ultimately, the macroscopic phenomenon of protein folding is itself the statistical result of an enormous number of microscopic atomic motions taking place within it. So could we actually calculate these microscopic movements, little by little? One of the major projects that took on this challenge is Folding@home.
Folding@home, launched at Stanford University in 2000 and still active today, is one of the world’s largest distributed computing network projects, connecting vast numbers of personal computers around the world. As its name suggests, it uses the computing resources of volunteers’ computers to carry out molecular dynamics (MD) simulations of proteins that would be difficult for a single computer to perform alone. The project runs large numbers of molecular dynamics simulations on the same protein, with the goal of calculating the movements of all the atoms that make up the protein and the surrounding water molecules, thereby reproducing the actual folding process inside a computer.
Distributed computing uses the computing resources of PCs around the world to handle an enormous amount of computation.
Molecular dynamics, which tracks how a protein folds and moves over time, applies Newton’s equations of motion and the laws governing physical forces at the molecular level. In a simulation, the movements of the atoms that make up the protein and its surrounding environment are calculated at extremely short time intervals—around 2 femtoseconds (\(2 \times 10^{-15}\) seconds). First, the forces acting between the atoms are calculated, and from those forces the positions to which the atoms will move at the next instant are determined. The forces are then calculated again at the new positions, followed by the next positions, and so on. By repeating these calculations continuously at extremely short intervals, we can track how individual atoms move and vibrate and how they approach and move away from one another over time, almost as if filming a movie. It is not a still image, but a process followed through time.
The problem is that the time interval required for these calculations is extraordinarily short. Because atoms move rapidly, properly tracking their movements at the atomic level requires their positions to be recalculated continuously in extremely fine time steps on the order of femtoseconds (\(10^{-15}\) seconds). [1] It is generally known that the folding of a protein in vivo often takes place on timescales ranging from microseconds to milliseconds. This means that reproducing the actual folding process of a protein at the atomic level requires an almost unimaginable amount of computation. To follow atomic motion for just 1 millisecond, calculations at femtosecond intervals must be repeated as many as 500 billion times.
\(1\text{ millisecond}(1\text{ ms}) = 10^{-3}\text{ seconds} = 0.001\text{ seconds}\), while \(2\text{ femtoseconds}(2\text{ fs}) = 2 \times 10^{-15}\text{ seconds} = 0.000000000000002\text{ seconds}\). If we calculate how many 2-femtosecond intervals are required to make up the target time of 1 millisecond:\[ \frac{10^{-3}\text{ seconds}}{2 \times 10^{-15}\text{ seconds}} = 500,000,000,000\text{ steps (500 billion)} \]
This means that to complete a movie covering the 0.001 seconds during which a protein folds, 500 billion snapshots taken at 2-femtosecond intervals would be needed in sequence. Moreover, the atomic positions at the second time step necessarily depend on the result of the first, so these calculations must proceed sequentially.
And that is not all. A single simulation is never enough. During one folding event, a protein can experience countless structures and pathways, and calculating only one trajectory cannot reveal the overall picture. Many simulations must therefore be repeated from different initial conditions, and the many trajectories obtained from them must be analysed together. This is precisely why the enormous distributed computing power of Folding@home was essential. Instead of using one supercomputer to make one long molecular movie, the calculations are divided among the computers of countless volunteers around the world, much like filming many short molecular movies simultaneously from different initial conditions. Each PC records the movements of the protein atoms over its assigned period as a trajectory, and the many independent trajectories collected in this way are then analysed statistically. Rather than filming one movie all at once, repeatedly filming the folding of the same protein from many different starting points allows us to compare not just the pathway taken during a single folding event, but the many different pathways and structures that appear across many folding events. In this way, researchers can discover different pathways and intermediate structures that would not be visible in a single folding event. By statistically analysing the enormous amount of collected data, they can also determine which structures the protein remains in for longer periods, which states it frequently moves between, which pathways it follows to reach the native state, and in which states folding is delayed. At that point, the energy landscape moves a significant step closer from being a metaphor in a diagram to becoming a physically calculable object.
Markov State Model (MSM)
Yet examining all of these trajectories directly is still far too complicated, because millions of atomic coordinates are continually changing over time. Researchers therefore simplify the data. Although the protein is constantly moving, a closer look reveals that similar structures that can move back and forth between one another within short periods appear repeatedly. Researchers group these similar structures, which can easily interconvert through thermal fluctuations, into a single state. They then analyse the many trajectories and calculate how frequently the protein moves from one state to another—in other words, the transition probability—thereby constructing a vast transition network. This is the Markov State Model, or MSM.
Rather than displaying every atomic movement of a protein directly, an MSM compresses the complex structural space into a number of states and represents protein dynamics using the transition probabilities and timescales between those states. This is quite similar to the way a navigation app finds the optimal route. [2] By analysing the GPS movement records sent by many drivers, we can determine how frequently cars appear on a particular road and how often they move from road A to road B. An MSM works in a similar way. By recording how often a protein remains in a particular structural region and how often it moves from one structure to another, these movements can be represented as a transition network.
Like the folding funnel, an MSM does not assume a single predetermined folding pathway. Instead, it draws the many states a protein can visit and the branching routes between them as a network. However, the folding funnel and an MSM show different things. The funnel model is a conceptual diagram showing how a protein as a whole explores its way towards the lower-free-energy native basin, whereas an MSM is a data-based map showing which states the protein actually passes through during that process and how it moves between them. In other words, if the funnel shows the overall landscape, the MSM shows how the protein actually moves across that landscape.
The figure below shows an example from an actual Folding@home study in which the folding of ACBP (acyl-coenzyme A-binding protein), a small protein consisting of 86 amino acids, was analysed using an MSM. [3] Using data obtained from numerous molecular dynamics simulations, the researchers divided the protein into 2,000 states and analysed the transitions occurring between them. The image does not show the entire complex network. Instead, it extracts 15 major folding pathways leading from the unfolded state to the native state. Rather than showing a typical MSM network in its complete form, this figure can be thought of as showing, through an MSM, which folding pathways a real protein mainly uses. The darker arrows reflect folding flux. Put simply, this value indicates how many folding events use a particular pathway as the protein moves from the unfolded state towards the native state. It is similar to many people travelling towards the same destination: some roads carry large numbers of vehicles, while others are used by very few. In this figure, therefore, pathways used by a greater number of folding events are shown with thicker arrows. The figure shows that a protein does not fold along one predetermined route from the beginning, but instead passes through many intermediate states and branching points. As the differences in arrow intensity indicate, not all pathways are used equally; some pathways carry a greater proportion of the folding process than others.
Drawing the energy landscape from transition probabilities, equilibrium probabilities, and the Boltzmann distribution
Using an MSM, we can learn how often and how quickly a protein moves from one state to another. But how can this information be connected to an energy landscape? This is where statistical physics enters the picture again. An MSM divides numerous molecular dynamics trajectories into different states and statistically analyses the transitions between them. The proportion occupied by each state in the equilibrium ensemble over a long period can be thought of as its equilibrium probability. Equilibrium probability means the proportion of time a protein spends in a particular state when observed over a sufficiently long period. For example, if, in a sufficiently sampled equilibrium system, a protein spends 30% of the total time in state A and 5% in state B, state A appears with a much higher probability than state B. This information provides an important clue for understanding the free-energy landscape. If we compare the free-energy landscape to a system of valleys, a protein remains longer in deep valleys and is less likely to remain for long periods in high regions. Thus, a state with a high equilibrium occupancy probability can be interpreted as a stable region with relatively low free energy, while a state with a low probability can be interpreted as a region with relatively high free energy.
Transition probability, on the other hand, provides somewhat different information. It shows how easily a protein moves from one state to another—in other words, it describes the protein’s dynamics. If transitions between two states occur very rapidly and frequently, the two states can be regarded as strongly connected dynamically. Conversely, if transitions rarely occur, moving between the two states is difficult. In terms of an energy landscape, this can be related to the energy barrier between two valleys. If the barrier is low, the protein can move between the two states relatively easily; if the barrier is high, the transition will be more difficult. However, it is important to note that the transition probability itself does not directly represent the height of the energy barrier. This is because actual transition rates are influenced not only by the energy barrier but also by factors such as solvent friction and the movements of surrounding molecules. Put very simply, equilibrium occupancy probability tells us how stable a state is, while transition probabilities and transition rates tell us how easily the protein moves between states.
From equilibrium probability to free energy: the Boltzmann distribution
If we know how long a protein remains in each state, then how can we turn that information into the highs and lows of free energy? This is where Boltzmann appears again. While studying gas molecules, Boltzmann showed that under thermal equilibrium there is a definite mathematical relationship between the energy of a particle and the probability that it will exist in a particular state. He mathematically explained that lower-energy states occur with higher probability, while the probability of higher-energy states falls rapidly. Because proteins are also in thermal equilibrium, the same principle can be applied. In other words, if a particular structure is frequently observed over a long period, it can be interpreted as a stable state with low free energy, while a structure that appears only rarely can be interpreted as an unstable state with high free energy. A high equilibrium probability means that the probability of being in that state is very high, indicating a stable structure with low free energy, whereas a low equilibrium probability can be interpreted as a state with relatively high free energy.
Mathematically, the equilibrium probability of a state and its free energy can be connected through the Boltzmann distribution. Under thermal equilibrium, the probability \(P_i\) of observing a state \(i\) is related to its free energy \(F_i\) as follows:\[ P_i \propto e^{-F_i / k_B T} \]
Rearranging this in terms of free energy gives:\[ F_i = -k_B T \ln P_i + C \]
Here, \(P_i\) is the equilibrium occupancy probability of state \(i\), \(F_i\) is the free energy corresponding to that state, \(k_B\) is the Boltzmann constant, \(T\) is the absolute temperature, and \(C\) is a constant that sets the reference point for free energy. The equation itself may look complicated, but the central idea is very simple. At equilibrium, states that are observed more frequently have lower free energy, while states observed less frequently have higher free energy. Therefore, if enough molecular dynamics simulations are performed to determine the proportion of total time occupied by each state, the resulting distribution can be used to estimate the relative free-energy differences between states. In other words, if we know how long a protein remains in a particular state, we can work backwards to estimate how deep that state lies on the free-energy landscape. Returning to the earlier example in which the protein spends 30% of the total time in state A and 5% in state B, state A can be interpreted as lying lower on the free-energy landscape—a deeper valley—than state B. Importantly, this does not mean that a state simply happened to be observed frequently at one particular moment. It refers to the probability occupied by that state within a sufficiently sampled equilibrium ensemble. In real proteins, rather than treating one specific structure as a single state, similar structures are often grouped together as one state, or basin, so the occupancy probability of that state also reflects the number and probabilities of the many microstates contained within it. Therefore, if we can determine sufficiently well how long a protein remains in each state, we can estimate how deep that state lies relative to others on the free-energy landscape.
Folding@home: starting from motion, obtaining probabilities, and inferring free energy from those probabilities
The meaning of large-scale molecular dynamics simulations such as Folding@home now becomes clearer. Molecular dynamics simulations calculate the positions and velocities of the atoms that make up a protein and its surrounding solvent at very short time intervals, producing long trajectories. Folding@home does not generate only one such trajectory; it produces large numbers of trajectories in parallel from different initial conditions. By collecting so many trajectories, researchers can statistically compare diverse structures and folding pathways that would be difficult to observe in a single simulation. This is also why Folding@home uses MSMs as an important method of analysis. When the data are organised into an MSM, we can obtain information such as how often the protein visits particular states, how quickly it moves between states, which states persist for longer periods, and which pathways it uses to move from one state to another. By applying equilibrium statistics and the Boltzmann relationship to this information, relative free-energy differences can be estimated from data obtained from actual molecular motion. In other words, the energy landscape can be inferred in reverse through the sequence: atomic motion → trajectory → states and transitions → equilibrium occupancy probability → free energy. The energy landscape is therefore not simply a picture imagined in our heads.
If the movements of real atoms are calculated in sufficiently large numbers and the results are analysed statistically, the stable states that emerge, the barriers between those states, and the structure of the folding pathways can be re-expressed in the language of free energy. Of course, this does not mean that every possible microstate of a real protein can be calculated without missing a single one to produce a complete free-energy map. That is practically impossible. Unlike a small molecule such as the alanine dipeptide discussed in the previous article, whose free-energy landscape can be represented using only the two dihedral angles φ and ψ, real proteins have far too many structural degrees of freedom, including numerous residues, bond angles, and torsion angles. The number of possible microstates is effectively astronomical, making it practically impossible to draw the entire free-energy landscape directly. Therefore, rather than calculating this enormous space one state at a time, actual research samples as much as possible of the regions that have been sufficiently explored through molecular dynamics and statistically analyses those data to reconstruct the important features of the free-energy landscape. This is also why the large-scale distributed computing of Folding@home is important. A single long trajectory cannot sufficiently reveal the complex structural space explored by a protein. By generating many short trajectories in parallel and linking them statistically, researchers can study slow dynamics and diverse pathways that are not visible in individual trajectories.
Microscopic statistics create the macroscopic landscape
Is the energy landscape really just a theoretical picture? Folding@home provides an important answer to this question. What we first calculate directly on a computer is not the picture of an energy landscape. What we calculate first is the movement of atoms. We record how the protein and surrounding atoms move over time in numerous trajectories and statistically analyse the structures and states that repeatedly appear, the transitions between states, and the equilibrium occupancy probability of each state. Only then do we connect those statistical results to free energy. In other words, we first calculate motion, then obtain probabilities, and finally infer free energy from those probabilities.
And this process is remarkably similar to the way Boltzmann thought when developing statistical mechanics more than 150 years ago. He showed that even if we cannot directly observe the movements of countless invisible atoms and molecules one by one, we can explain the macroscopic phenomena we observe with our own eyes if we know the statistical distribution of that microscopic world. In thermodynamics, the statistical results of the microscopic system appear as concepts such as entropy and free energy. In proteins, the same statistical principles appear in a much more complex form. By analysing the statistical distributions produced by the motions of countless atoms, we can understand which structures are stable, which states persist for long periods, which states readily interconvert, and which pathways lead towards the native basin. By connecting this information to free energy, features of the energy landscape that we had previously thought of theoretically—such as basins, energy barriers, and kinetic traps—can be estimated from actual computational data. Seen this way, the energy landscape is not simply a metaphor. Of course, we cannot draw the entire enormous, high-dimensional free-energy landscape of a real protein exactly as it exists. The diagrams we see are simplified maps produced by projecting a complex high-dimensional space onto particular reaction coordinates or structural variables. Even so, the important features of that landscape can be inferred and tested through actual molecular motion and statistical analysis. Energy landscape theory is fascinating precisely because the valleys and barriers that began as theoretical pictures can reappear through data obtained from real molecular movements and statistical mechanics.
Folding@home and AlphaFold: complementary technologies
These calculations and the accumulation of data are still continuing today. Folding@home produces large-scale molecular dynamics trajectories and makes available data that can be used to study protein structural ensembles, dynamics, MSMs, and related phenomena. Meanwhile, on another front in the field of protein folding, artificial-intelligence-based technologies such as AlphaFold are predicting protein structures. These systems learn structural patterns and correlations that repeatedly appear across vast amounts of protein sequence and structure data and, given an amino acid sequence, predict which three-dimensional structure that sequence is most likely to adopt. Rather than being competing technologies, the two address different areas and can complement one another. If AlphaFold’s strength lies in predicting what structure a sequence is likely to form, molecular dynamics and Folding@home are powerful tools for physically investigating how that structure forms, which intermediate states it passes through, and how quickly it moves between different states. In actual research, the boundary between the two is not clearly divided, and these computational approaches, with their different strengths, may become even more closely integrated in the future. Ultimately, understanding a protein does not end with identifying a single final three-dimensional structure. We must also consider the countless microstates that give rise to that structure and their probabilities, the transitions and dynamics between different states, and the continual competition and balance between enthalpy and entropy occurring within the surrounding environment that includes both the protein and the solvent. In the end, all of these connect into a single picture. The macroscopic phenomena we can see—protein folding and protein function—arise from the invisible movements of countless atoms and their statistical distributions. Folding@home is a fascinating example of showing this connection through actual computation. Seen in this way, the funnel we have examined is not simply a diagram. It is a map that compresses into a single image the macroscopic order statistically produced by movements in the microscopic world.
[References]
[3] Slow unfolded-state structuring in ACBP folding revealed by simulation and experiment
[Image Sources]
Image 1 Energy landscape theory of protein folding — CC BY 4.0
Image 2 ACBP MSM from Folding@home — CC BY-SA 3.0

